Papers with American English

34 papers
Predicting the Target Word of Game-playing Conversations using a Low-Rank Dialect Adapter for Decoder Models (2025.naacl-short)

Copied to clipboard

Challenge: Existing work proposes dialect adaptation for encoder models or encoder-decoder models.
Approach: They propose to use MD-3 to combine task adapters and dialect adapters to decoder models using a masked word game-playing conversation.
Outcome: The proposed architecture outperforms baselines on Indian English and Nigerian English on a masked conversation with two models.
“Talk to me with left, right, and angles”: Lexical entrainment in spoken Hebrew dialogue (2021.eacl-main)

Copied to clipboard

Challenge: entrainment is a widespread phenomenon in human interaction that leads interlocutors to adapt their linguistic productions to become more similar to each other.
Approach: They propose to use existing measures to analyze Hebrew speakers interacting in a Map Task to find evidence of lexical entrainment.
Outcome: The proposed study is the first to examine lexical entrainment in Hebrew using two existing measures.
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties (2024.eacl-long)

Copied to clipboard

Challenge: Mainstream cross-lingual task-oriented dialogue systems often overlook the transfer to lower-resource colloquial varieties due to limited test data.
Approach: They propose to train a model for intent recognition and slot-filling in English and apply it to other languages.
Outcome: The proposed model performs better than existing models on English and other languages.
Multi-VALUE: A Framework for Cross-Dialectal English NLP (2023.acl-long)

Copied to clipboard

Challenge: Current systems that focus on standard American English are not dialect invariant . current systems focus on a single dialect, which results in performance discrepancies .
Approach: They propose a resource for evaluating and achieving English dialect invariance . they stress test question answering, machine translation, and semantic parsing .
Outcome: The proposed system is based on a rule-based translation system spanning 50 English dialects and 189 unique linguistic features.
TADA : Task Agnostic Dialect Adapters for English (2023.findings-acl)

Copied to clipboard

Challenge: Existing work on dialectal English NLP is task-specific, using manual annotated dialect data, weak supervision, or data augmentation.
Approach: They propose a method for task-agnostic dialect adaptation by aligning non-SAE dialects with task-specific adapters from SAE.
Outcome: The proposed method improves dialectal robustness on 4 dialectal variants of the GLUE benchmark without task-specific supervision.
“Cheese!”: a Corpus of Face-to-face French Interactions. A Case Study for Analyzing Smiling and Conversational Humor (2020.lrec-1)

Copied to clipboard

Challenge: Cheese! is a conversational corpus containing 11 mixed and non-mixed dyadic interactions lasting around 15 minutes each.
Approach: They propose to use a conversational corpus to compare smiling behavior in American English and French conversations to conduct a cross-cultural comparison.
Outcome: The proposed study examines the relationship between smile and humor in conversational interactions between American English and French participants.
Meaning Variation and Data Quality in the Corpus of Founding Era American English (2025.acl-short)

Copied to clipboard

Challenge: Legal scholars are increasingly using corpus based methods for assessing historical meaning . main corpus used in legal arguments is the Corpus of Founding Era American English .
Approach: They demonstrate how NLP can be used to infer meaning change and variation using masked language models.
Outcome: The proposed method can be used to infer meaning change and variation using advanced methods.
A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American English (2020.coling-main)

Copied to clipboard

Challenge: Despite the high performance of transformer-based language models, we still lack understanding of the kind of linguistic knowledge they learn and rely on.
Approach: They evaluate three transformer-based language models and test their grammatical and semantic knowledge by sentence-level probing, diagnostic cases, and masked prediction tasks.
Outcome: The models capture grammatical and semantic knowledge, but they lack model-specific weaknesses especially on semantic knowledge.
SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition (P19-1)

Copied to clipboard

Challenge: Selectional Preference (SP) is a common phenomenon in human language and has been shown to be useful in many natural language processing tasks.
Approach: They propose a large-scale evaluation set that provides human ratings for the plausibility of 10,000 SP pairs over five SP relations, covering 2,500 most frequent verbs, nouns, and adjectives in American English.
Outcome: The proposed evaluation sets provide human ratings for plausibility of 10,000 SP pairs over five SP relations covering 2,500 most frequent verbs, nouns, and adjectives in American English.
Spelling convention sensitivity in neural language models (2023.findings-eacl)

Copied to clipboard

Challenge: Various long-distance dependencies have been investigated using neural language models.
Approach: They examine whether large neural language models learn the long-distance dependency of British versus American spelling conventions . a large T5 language model does internalize consistency, but only with respect to observed lexical items .
Outcome: The proposed model internalizes consistency with the training corpora, but only with respect to observed lexical items.
DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects (2026.acl-long)

Copied to clipboard

Challenge: Current disinformation detection systems are predominantly developed and evaluated on Standard American English (SAE) . however, their robustness to dialectal variation is unexplored.
Approach: They propose a benchmark for evaluating disinformation detection robustness across 50 English dialects . they use multi-value's linguistically-grounded transformations to introduce D-CUBE (Dialectal Disinformation Detection Corpus)
Outcome: The proposed model outperforms zero-shot LLMs in human-written dialects while AI-generated content remains stable.
The Risk of Racial Bias in Hate Speech Detection (P19-1)

Copied to clipboard

Challenge: Annotators’ insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations.
Approach: They propose *dialect* and *race priming* as ways to reduce the racial bias in hate speech detection models by detecting differences in dialects in annotated tweets.
Outcome: The proposed models acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others.
Perceptions of Language Technology Failures from South Asian English Speakers (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies have identified performance disparities between Standard American English and other English dialects, but the degree to which these discrepancies affect user experience is not well understood.
Approach: They aim to reduce performance gap for South Asian Englishes by surveying their interactions with language technology and comparing their results to a control survey.
Outcome: The proposed model reduces the performance gap for South Asian Englishes (SAsE) speakers are more likely to recall failures with language technology and to reference specific issues with written language technology than SAmE speakers.
VALUE: Understanding Dialect Disparity in NLU (2022.acl-long)

Copied to clipboard

Challenge: English Natural Language Understanding systems outperform humans on benchmarks like GLUE and SuperGLUE, but they only use textbook Standard American English (SAE) . fewer studies have considered the effects of dialectal differences on performance .
Approach: They propose a benchmark to evaluate the performance of English natural language understanding systems using a set of lexical and morphosyntactic transformation rules.
Outcome: The proposed model outperforms humans on GLUE and SuperGLUE, but only on standard American English . the proposed model recruits fluent speakers of African American vernacular english to validate each feature transformation .
My LLM might Mimic AAE - But When Should It? (2025.naacl-long)

Copied to clipboard

Challenge: a study examines the representation of African American English in large language models . a survey of black americans and annotation of LLM outputs shows that Black Americans prefer to use AAE in formal settings .
Approach: They examine Black Americans' perceptions of how effective AI tools are at producing authentic African American English in large language models.
Outcome: The results show that Black Americans prefer to use LLMs in formal settings over informal ones . the results show they prefer to produce AAE in less formal settings .
Challenges in Automated Debiasing for Toxic Language Detection (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for debiasing toxic language data are limited in their ability to prevent biased behavior in toxic language detection systems.
Approach: They propose to debiase toxic language detection models using lexical and dialectal markers using synthetic labels instead of traditional methods.
Outcome: The proposed method reduces dialectal associations with toxicity despite the use of synthetic labels .
Towards Building an Automatic Transcription System for Language Documentation: Experiences from Muyu (2020.lrec-1)

Copied to clipboard

Challenge: Language documentation is a rapidly growing field due to its urgency.
Approach: They propose to use phoneme recognition to automatically recognize spoken languages and translate them to global languages.
Outcome: The proposed tool performs better than existing methods with American English, Austrian German and Slovenian as source and target languages.
Cross-Cultural Transfer Learning for Text Classification (D19-1)

Copied to clipboard

Challenge: a large dataset is required to achieve competitive performance in most natural language tasks. large datasets are expensive, time consuming, and error-prone.
Approach: They propose a transfer-learning framework that leverages bilingual corpora for natural language text classification using no task-specific data.
Outcome: The proposed framework can achieve good performance on formality classification and sarcasm detection tasks without any task-specific labeled data.
Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection (2022.naacl-main)

Copied to clipboard

Challenge: toxicity annotations are often ignored because of its subjective nature and lack of nuance.
Approach: They examine the effect of annotator identities and beliefs on toxic language annotations by considering posts with three characteristics: anti-Black language, African American English (AAE) dialect, and vulgarity.
Outcome: The findings show strong associations between annotator identity and beliefs and ratings of toxicity.
Analysis of LLM as a grammatical feature tagger for African American English (2025.findings-naacl)

Copied to clipboard

Challenge: African American English (AAE) presents unique challenges in natural language processing (NLP).
Approach: They evaluate the ability of different NLP systems to recognize distinctive AAE grammatical features by using sentence-level binary classification tasks using both zero-shot and fewshot strategies.
Outcome: The evaluation involved sentence-level binary classification tasks, using both zero-shot and few-shot strategies.
The Niki and Julie Corpus: Collaborative Multimodal Dialogues between Humans, Robots, and Virtual Agents (L18-1)

Copied to clipboard

Challenge: Niki and Julie corpus contains more than 600 dialogues between humans and robots . corpus includes audio and video recordings, results of ranking tasks, questionnaire responses .
Approach: the corpus contains more than 600 dialogues between human participants and a robot . the dialogues are part of a collaborative item-ranking task designed to measure influence .
Outcome: the corpus contains more than 600 dialogues between human participants and a robot or virtual agent . the dialogues contain conversational errors by the robot, which simulates typical of modern automated agents .
Investigating African-American Vernacular English in Transformer-Based Text Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work in Natural Language Generation (NLG) uses a Transformer-based language model to generate high-quality, coherent text when prompted by arbitrary input.
Approach: They evaluate the performance of a Transformer-based model that generates high-quality, coherent text when prompted by arbitrary input.
Outcome: The proposed model improves on AAVE and SAE text with pretrained sentiment classifiers.
Task-Agnostic Low-Rank Adapters for Unseen English Dialects (2023.emnlp-main)

Copied to clipboard

Challenge: a recent study found that LLMs are trained on corpora disproportionally weighted in favor of Standard American English . prior work on dialect struggle with generalizing to evolving and emerging dialects in a scalable manner.
Approach: They propose a method that leverages linguistic knowledge to enable resource-efficient adaptation . their method disentangles dialect-specific and cross-dialectal information .
Outcome: a new method improves generalization to unseen dialects in a task-agnostic fashion . it achieves the best or most competitive performance across 5 dialects .
A Model of the Language Process (2026.acl-long)

Copied to clipboard

Challenge: Language is a process that changes over time as new vocabulary emerges, word meanings shift, and narratives progress.
Approach: They introduce a BERT style transformer encoder that models language by jointly learning to predict document contents and classify document publication dates.
Outcome: The proposed model can predict document contents and classify document publication dates and accurately detects changes in word meanings.
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination (2024.emnlp-main)

Copied to clipboard

Challenge: a large-scale study of linguistic bias exhibited by ChatGPT covers 10 dialects of English . standard varieties of English, especially SAE, dominate available training data .
Approach: They use ChatGPT to generate models that default to "standard" varieties of English . they also use a feature annotation and native speaker evaluation to analyze the responses .
Outcome: The proposed models default to "standard" varieties of English, but non-"standard" ones exhibit stereotyping, demeaning content, lack of comprehension, condescending responses.
DADA: Dialect Adaptation via Dynamic Aggregation of Linguistic Rules (2023.emnlp-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) that focus on Standard American English (SAE) often suffer from performance degradation when applied to other dialects.
Approach: They propose a modular approach to imbue SAE-trained models with multi-dialectal robustness . they propose adapters which handle specific linguistic features to imbibe SAe-taught models .
Outcome: The proposed approach improves performance across multiple dialects and dialects.
Language Variety Identification with True Labels (2024.lrec-main)

Copied to clipboard

Challenge: Language identification datasets are compiled with the assumption that the gold label of each instance is determined by where texts are retrieved from.
Approach: They present a human-annotated multilingual dataset for language variety identification . they use a model to train multiple models to discriminate between different languages .
Outcome: The proposed dataset provides a reliable benchmark toward robust and fairer language variety identification systems.
Modeling Gender and Dialect Bias in Automatic Speech Recognition (2024.findings-emnlp)

Copied to clipboard

Challenge: Dialect and gender-based biases have become an area of concern in language-dependent AI systems.
Approach: They construct a podcast audio dataset and evaluate its performance . they then refine the models to better understand how finetuning can impact performance.
Outcome: The proposed model improves on 13 hours of podcast audio transcribed by speakers of four US-based English dialects.
Leveraging Syntactic Dependencies in Disambiguation: The Case of African American English (2024.lrec-main)

Copied to clipboard

Challenge: African American English (AAE) is a low-resource language facing the challenge of inadequate annotated data for training natural language processing models.
Approach: They propose a syntactically informed classifier for automatic disambiguation of AAE's habitual be.
Outcome: The proposed classifier improves automatic disambiguation of habitual and non-habitual meanings of "be" integrating syntactic information improves disambiguations of habituality by 65 F1 points over baseline models and as much as 74 points.
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks often overlook intra-language variations, leaving speakers of non-standard dialects underserved.
Approach: EnDive evaluates seven state-of-the-art large language models across tasks . human evaluations confirm high translation quality, with average scores of at least 6.02/7 .
Outcome: EnDive evaluates state-of-the-art large language models across language understanding, reasoning, mathematics, logic tasks.
NarrativeTime: Dense Temporal Annotation on a Timeline (2024.lrec-main)

Copied to clipboard

Challenge: e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text.
Approach: They propose a timeline-based framework that achieves full coverage of all possible TLINKs.
Outcome: The proposed framework achieves full coverage of all possible TLINKs in a text.
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies show voice assistants do not perform equally well for everyone . however, research on demographic robustness of speech technologies is still scarce .
Approach: They propose a statistical method to detect demographic bias using a large dataset with controlled demographic tags.
Outcome: The proposed method shows statistically significant differences in performance across age, dialectal region and ethnicity.
A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems (2025.acl-long)

Copied to clipboard

Challenge: Existing Privacy Policy Question Answering systems exhibit performance disparities across English dialects, disadvantaging speakers of non-standard varieties.
Approach: They propose a framework that integrates a Dialect Agent and a Privacy Policy Agent to mitigate dialectal biases.
Outcome: The proposed framework improves GPT-4o-mini’s zero-shot accuracy from 0.394 to 0.601 on PrivacyQA and 0.352 to 0.464 on PolicyQA.
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations (2026.acl-long)

Copied to clipboard

Challenge: Agentic benchmarks rely on LLM-simulated users to evaluate agent performance . however, the robustness, validity, and fairness of this approach remain unexamined .
Approach: They investigate whether LLM-simulated users are reliable proxies for real human users . they find that agent success rates vary up to 9 percentage points across different LLMs .
Outcome: The results show that simulated users underestimate success on challenging tasks while miscalibrate performance on moderately difficult tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations